Back

Journal of Speech, Language, and Hearing Research

American Speech Language Hearing Association

All preprints, ranked by how well they match Journal of Speech, Language, and Hearing Research's content profile, based on 13 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Late-Talking Children Talk More? A Machine Learning Approach to Speech Act Analysis in Early Childhood

Dhakal, G.; He, H.; Newman, S. D.; Xiong, Y.

2025-10-02 pathology 10.1101/2025.09.30.679667 medRxiv
Top 0.1%
47.4%
Show abstract

Speech acts shape early language development and social cognition, yet little is known about how late-talking (LT) children use them to achieve communicative goals. We compared LT and typically developing (TD) preschoolers (1;09-6;00) across nine dyadic English corpora, using a Conditional Random Field model to annotate speech acts. We analyzed speech act distributions, hierarchical relations, and contingent responses to assess production and comprehension. TD children produced more declarative statements and wh-questions, whereas LT children produced more unclear word-like utterances and showed reduced comprehension ability (1;09-2;07). Speech acts classified LT and TD groups with 72.3% accuracy, improving to 76.6% with linguistic and demographic features. Classification was driven by co-occurring patterns of speech act frequencies. LT children showed delayed onset of speech acts but employed more speech acts after 3;09, focusing on speaker-centered goals, whereas TD children favored collaborative use of speech acts, revealing complex dynamics in the development of communicative skills.

2
A Meta-Analytical Review of Executive Function Skills in Adults who Stutter

Ofoe, L. C.; Ntourou, K.; Clifton, S.; Coalson, G. A.

2025-09-04 pathology 10.1101/2025.09.02.25334917 medRxiv
Top 0.1%
45.0%
Show abstract

PurposeExecutive function has been identified as a potential area of vulnerability in individuals who stutter. The present study identified and analyzed the data across empirical studies of the executive function skills of adults who do (AWS) and do not stutter (AWNS). MethodElectronic databases, literature reviews, and reference sections of articles and dissertations were searched to identify candidate studies that examined behavioral measures of working memory, inhibition, and/or cognitive flexibility. A total of 39 studies met the eligibility criteria for this meta-analysis. A random-effects model was applied to estimate the pooled effect sizes (Hedges g) and 95% confidence intervals. ResultsAWS were significantly less accurate than AWNS on measures of working memory (Hedges g = - 0.41, p < .001), including nonword repetition (Hedges g = -.57, p < .001), forward digit span (Hedges g = -0.23, p = .02), backward digit span (Hedges g = -.38, p = .004) and operation span tasks (Hedges g = - .37, p = .018). AWS performed comparably to AWNS on inhibition measures (Hedges g = -0.10, p = .37). An insufficient number of published studies were available to conduct a meaningful analysis of cognitive flexibility. ConclusionsPresent findings suggest that AWS, as a group, exhibit weaknesses in one component of executive function - working memory - compared to AWNS. Additional research is necessary to determine potential differences in inhibition and cognitive flexibility in AWS.

3
Self-reported effects of classic psychedelics on stuttering

Gold, N.; Goldway, N.; Gerlach-Houck, H.; Jackson, E. S.

2023-04-20 pathology 10.1101/2023.04.18.537312 medRxiv
Top 0.1%
25.6%
Show abstract

Stuttering is a neurodevelopmental communication disorder that can lead to significant social, occupational, and educational challenges. Traditional behavioral interventions for stuttering can be helpful, but effects are often limited. Classic psychedelics hold promise as a complement to traditional interventions, but their impact on stuttering is unknown. We conducted a qualitative content analysis to explore potential benefits and negative effects of psychedelics on stuttering using publicly available Reddit posts. A combined inductive-deductive approach was used whereby meaningful units were extracted and codes were initially assigned inductively. We then deductively applied an established framework to organize the effects (i.e., codes) into five subthemes (Behavioral, Emotional, Cognitive, Belief, and Social Connection), each of which was grouped under an organizing theme (positive, negative, neutral). Results indicated that the effects of psychedelics spanned all subthemes. Nearly 75% of participants reported overall positive effects. Nearly 60% of participants indicated positive behavioral change (e.g., reduced stuttering, increased speech control), 40% reported positive emotional benefit, 15% reported positive cognitive changes, 12% reported positive effects on beliefs, and 7% indicated positive social effects. Approximately 10% of participants reported negative behavioral effects (e.g., increased stuttering, reduced speech control). Psychedelics may help many stutterers improve communication, cultivate a healthier outlook, and promote psychological well-being. These preliminary results indicate that future clinical trials investigating psychedelic-assisted speech therapy for stuttering are warranted.

4
Evaluating Goodness of Pronunciation and Phonological Posteriors as Objective Markers of Speech Severity in Motor Speech Disorders

Wang, F.; Utianski, R. L.; Duffy, J. R.; Barnard, L. R.; Botha, H.

2026-07-16 neurology 10.64898/2026.07.14.26358076 medRxiv
Top 0.1%
23.3%
Show abstract

This study examined the extent to which goodness of pronunciation (GoP) scores and phonological posterior probabilities capture perceptual ratings of speech severity in individuals with motor speech disorders (MSD). Speech recordings of the word catastrophe were obtained from 489 participants, including 333 neurologically typical controls and 156 individuals with MSD. GoP scores were derived using traditional acoustic features and self-supervised speech representations, including WavLM and XLS-R, across multiple modeling approaches, while phonological posterior probabilities were extracted using Phonet. Model performance was evaluated using Kendall's rank correlations, regression, and receiver operating characteristic analyses against speech-language pathologists' perceptual ratings of sound distortion and intelligibility. Both GoP and phonological posterior probabilities were significantly associated with perceptual ratings. Self-supervised speech representations substantially outperformed traditional acoustic features, with WavLM-based GoP using k-nearest neighbors achieving the strongest performance. Across correlation, regression, and classification analyses, GoP consistently outperformed phonological posterior probabilities for both sound distortion and intelligibility. Age and gender had minimal influence on model-derived measures or their relationships with perceptual ratings. These findings demonstrate the value of self-supervised GoP as an objective measure of speech impairment while highlighting the complementary role of phonological posterior probabilities in characterizing articulatory aspects of motor speech disorders.

5
EEG responses to auditory cues predict fluency variability and stuttering intervention outcome

Rocha, M. F.; Carmona, J.; Correia, J. M.

2025-02-24 neuroscience 10.1101/2025.02.21.635719 medRxiv
Top 0.1%
18.7%
Show abstract

Stuttering is a variable speech disorder whose brain mechanisms remain unknown. Sensorimotor brain circuits, critical for motor-speech control, including auditory processing necessary for speech prediction and monitoring, have been linked to the disorder. Despite considerable advances, it remains unclear whether auditory circuits relate to stuttering variability, and whether the panoply of interventions for persons who stutter can lead to brain changes within these circuits. We employed electroencephalography (EEG), in a group of persons who stutter, in combination with auditory probes to tap onto the importance of auditory cortical regions in stuttering variability. Participantsproduced flexible speech (i.e., describing visual scenes) and non-flexible speech (i.e., reading syllables), following an auditory cue. More pronounced P200 auditory evoked potentials were observed in participants with higher dysfluency rates, mainly in the spontaneous speech task. Interestingly, speech therapy intervention led to a reduction of the P200 potential, which was in turn significantly related to fluency improvements. Furthermore, EEG response patterns discriminative of cue frequency (400 or 800 Hz tones) were also predictive of dysfluency scores. Our study highlights the involvement of auditory cortical processing and that of auditory attention in stuttering variability. We support that a higher state of auditory alertness may be implicated in the sensorimotor mechanisms of stuttering, and that speech therapy interventions promoting more self-confident communication can restraint auditory alertness, and potentially reduce speech dysfluencies. HighlightsO_LIAuditory probes can assess the auditory cortex in speech production and stuttering. C_LIO_LIStuttering severity correlates to EEG auditory responses during speech preparation. C_LIO_LIHigher states of auditory alertness in stuttering may be reduced by speech therapy. C_LI Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=123 SRC="FIGDIR/small/635719v1_ufig1.gif" ALT="Figure 1"> View larger version (25K): org.highwire.dtl.DTLVardef@18a639dorg.highwire.dtl.DTLVardef@91efc6org.highwire.dtl.DTLVardef@114dd97org.highwire.dtl.DTLVardef@e01024_HPS_FORMAT_FIGEXP M_FIG C_FIG Speaking requires orchestrating several brain processes at a time. The auditory system assumes a central role, not only in waiting for the right moment to initiate speech, listening to self-produced speech, predicting the consequence of future speech, but also adjusting these processes to the intermittent nature of stuttering.

6
Iconic Sound-Shape Correspondences in Aphasia

Dorsi, J.; Sandberg, C.; Lacey, S.; Nygaard, L.; Sathian, K.

2026-05-19 neuroscience 10.64898/2026.05.18.725976 medRxiv
Top 0.1%
16.6%
Show abstract

PurposeTo examine speech iconicity for shape in aphasia, we compared iconicity ratings from people with aphasia to those from neurologically intact individuals and evaluated how iconicity relates to phonological and semantic processing profiles in aphasia. MethodEleven people with aphasia and 11 age- and gender-matched neurologically intact participants rated how rounded or pointed 50 auditory pseudowords sounded using a 5-point scale. Ratings from participants with aphasia were compared to predicted iconicity ratings derived from reference ratings from prior work and to ratings from neurologically intact participants. For each participant with aphasia, correlations between individual ratings and predicted ratings were related to measures of phonological and semantic processing. ResultsRatings from people with aphasia were significantly correlated with both the predicted ratings and the ratings from neurologically intact participants. The strength of the correlation between individual ratings and predicted ratings did not differ significantly between groups, although there was a trend toward weaker correlations in the aphasia group. There were indications that greater language impairment was associated with greater disruption of iconicity ratings; in particular, deficits in phonological segmentation and semantic processing were associated with reduced sensitivity to shape iconicity. ConclusionThese findings suggest that sensitivity to shape iconicity is preserved in individuals with aphasia to varying degrees. The specific nature of language impairment appears to play an important role in determining iconicity processing in aphasia.

7
Phonemic awareness deficits in an alphasyllabary language: Effects of task type and linguistic complexity in children with Specific Learning Disorder-Reading

Soman, A.; Dev, S. S.; Ravindren, R.

2026-04-07 psychiatry and clinical psychology 10.64898/2026.04.02.26349894 medRxiv
Top 0.1%
15.6%
Show abstract

Background Phonemic awareness deficits are a core feature of Specific Learning Disorder-Reading (SLD-R). How task- and language-specific factors influence these deficits in alphasyllabary languages may help clarify the cognitive mechanisms underlying reading impairment in SLD-R. Methods Thirty children with a DSM-5 diagnosis of SLD-R (mean age 11.4 years) and 29 age-matched typically developing children were given phoneme blending (words and pseudowords) and segmentation tasks in Malayalam. The effects of age and consonant clusters on task performance were evaluated. Results Children with SLD-R performed significantly worse than controls across most phonemic awareness tasks, with the largest deficits observed in pseudoword blending and word blending, and smaller deficits in segmentation. No significant difference was observed for initial phoneme deletion. In typically developing children, age showed strong positive correlations with phonemic performance across most tasks, whereas the SLD-R group showed weak or absent correlations, except in word blending and initial phoneme deletion. Consonant clusters significantly affected performance in both groups, with SLD-R showing more severe deficits. Conclusions Phonemic awareness deficits observed in SLD-R in alphasyllabary languages like Malayalam are more prominent in tasks where lexical support is absent, like pseudoword blending. These deficits vary across task types and linguistic complexity. Phonemic awareness improves with age in typically developing children, while improvement is uneven in children with SLD-R. The findings suggest that phonemic awareness deficits are a core feature of SLD-R across languages, but their manifestation is shaped by orthographic and linguistic characteristics of the writing system.

8
Does over-reliance on auditory feedback cause disfluency? An fMRI study of induced fluency in people who stutter.

Meekings, S.; Jasmin, K.; Lima, C. F.; Scott, S. K.

2020-11-20 neuroscience 10.1101/2020.11.18.378265 medRxiv
Top 0.1%
15.1%
Show abstract

This study tested the idea that stuttering is caused by over-reliance on auditory feedback. The theory is motivated by the observation that many fluency-inducing situations, such as synchronised speech and masked speech, alter or obscure the talkers feedback. Typical speakers show speaking-induced suppression of neural activation in superior temporal gyrus (STG) during self-produced vocalisation, compared to listening to recorded speech. If people who stutter over-attend to auditory feedback, they may lack this suppression response. In a 1.5T fMRI scanner, people who stutter spoke in synchrony with an experimenter, in synchrony with a recording, on their own, in noise, listened to the experimenter speaking and read silently. Behavioural testing outside the scanner demonstrated that synchronising with another talker resulted in a marked increase in fluency regardless of baseline stuttering severity. In the scanner, participants stuttered most when they spoke alone, and least when they synchronised with a live talker. There was no reduction in STG activity in the Speak Alone condition, when participants stuttered most. There was also strong activity in STG in response to the two synchronised speech conditions, when participants stuttered least, suggesting that either stuttering does not result from over-reliance on feedback, or that the STG activation seen here does not reflect speech feedback monitoring. We discuss this result with reference to neural responses seen in the typical population.

9
AI-based Speech Error Detection to Differentiate Primary Progressive Aphasia Variants

Vonk, J. M. J.; Lian, J.; Cho, C. J.; Antonicelli, G.; Ezzes, Z.; Wauters, L. D.; Keegan-Rodewald, W.; Kurteff, G. L.; Rodriguez, D. A.; Dronkers, N.; Henry, M. L.; Miller, Z. A.; Mandelli, M. L.; Anumanchipalli, G. K.; Gorno-Tempini, M. L.

2026-02-24 neurology 10.64898/2026.02.23.26346899 medRxiv
Top 0.1%
12.6%
Show abstract

BackgroundArtificial Intelligence (AI) based approaches to speech analysis have the potential to assist with objective speech error analysis in aphasia but off-the shelf tools often fail to detect speech errors due to prioritizing "fluent transcription." Speech production errors (dysfluencies) are hallmark diagnostic features of the nonfluent (nfvPPA) and logopenic (lvPPA) variants of primary progressive aphasia, yet they can be challenging to detect and characterize even by expert clinicians. This study aimed to evaluate whether the novel automated lightweight Scalable Speech Dysfluency Modeling system (SSDM-L), specifically designed to detect dysfluencies, could accurately distinguish PPA variants using voice recordings of individuals reading a brief passage. MethodParticipants included a total of 104 individuals, 40 with nfvPPA, 40 with lvPPA (matched on disease severity), and 24 healthy controls who read aloud the Grandfather Passage as part of a widely used motor speech evaluation (MSE). We automatically extracted ten speech error (dysfluency) variables using SSDM-L, including insertions, replacements, and deletions at both phoneme- and word-levels, and phoneme-level prolongations and repetitions. Group differences were assessed via ANCOVAs controlling for age, education, and disease severity (MMSE, CDR sum-of-boxes). To test clinical relevance, we performed correlation analyses with MSE ratings provided by experienced speech-language pathologists (i.e., gold standard) within the nfvPPA group. Classification performance was assessed by training random forest and XGBoost machine-learning models including 5-fold cross-validation. ResultsAll individuals read the entire passage in less than five minutes. SSDM-L detected eight of the ten predefined dysfluency features at sufficient frequency to include them in subsequent analyses. All eight features distinguished PPA from controls (p<.006). Individuals with nfvPPA made more errors than the lvPPA group on every feature (all p<.023). Each feature showed a moderate positive correlation with a global MSE apraxia/dysarthria score (r=.31-.56; p<.001-.053). Together, the eight features were able to classify nfvPPA versus lvPPA at AUC=.806 (random forest) and AUC=.776 (XGBoost). DiscussionAI-based automated speech error analysis accurately distinguished nfvPPA and lvPPA variants using a brief reading task. This quick error-sensitive scalable AI system has the potential of providing a practical tool to aid diagnosis in aphasia and motor speech disorders.

10
Functional Roles of Sensorimotor Alpha and Beta Oscillations in Overt Speech Production

Huang, L. Z.; Cao, Y.; Janse, E.; Piai, V.

2024-10-08 neuroscience 10.1101/2024.09.04.611312 medRxiv
Top 0.1%
12.4%
Show abstract

Power decreases, or desynchronization, of sensorimotor alpha and beta oscillations (i.e., alpha and beta ERD) have long been considered as indices of sensorimotor control in overt speech production. However, their specific functional roles are not well understood. Hence, we first conducted a systematic review to investigate how these two oscillations are modulated by speech motor tasks in typically fluent speakers (TFS) and in persons who stutter (PWS). Eleven EEG/MEG papers with source localization were included in our systematic review. The results revealed consistent alpha and beta ERD in the sensorimotor cortex of TFS and PWS. Furthermore, the results suggested that sensorimotor alpha and beta ERD may be functionally dissociable, with alpha related to (somato-)sensory feedback processing during articulation and beta related to motor processes throughout planning and articulation. To (partly) test this hypothesis of a potential functional dissociation between alpha and beta ERD, we then analyzed existing intracranial electroencephalography (iEEG) data from the primary somatosensory cortex (S1) of picture naming. We found moderate evidence for alpha, but not beta, ERDs sensitivity to speech movements in S1, lending supporting evidence for the functional dissociation hypothesis identified by the systematic review.

11
British Version of the Iowa Test of Consonant Perception

Guo, X.; Benzaquen, E.; Holmes, E.; Choi, I.; McMurray, B.; Bamiou, D.-E.; Berger, J. I.; Griffiths, T. D.

2024-09-07 neuroscience 10.1101/2024.09.04.611204 medRxiv
Top 0.1%
12.3%
Show abstract

The Iowa Test of Consonant Perception (ITCP) is a single-word closed-set speech- in-noise test with well-balanced phonetic features that provides a reliable testing option for real-world listening. Objectives. The current study aimed to establish a UK version of the test (B-ITCP) based on the British received pronunciation. Design. We conducted a validity test with 46 participants using the B-ITCP test, a sentence-in- noise test, and audiogram. Results. The B-ITCP demonstrated excellent test-retest reliability, cross-talker validity, and good convergent validity, consistent with the US results. Conclusions. These findings suggest that B-ITCP is a reliable measure of speech-in-noise perception, to facilitate comparative or combined studies in USA and UK. All materials (application and scripts) to run or construct the B-ITCP and ITCP are freely available online.

12
The impact of cognitive ability on multitalker speech perception in neurodivergent individuals

Lau, B. K.; Emmons, K.; Maddox, R. K.; Estes, A.; Dager, S.; (Astley) Hemingway, S.; Lee, A. K.

2022-09-20 psychiatry and clinical psychology 10.1101/2022.09.19.22280007 medRxiv
Top 0.1%
12.1%
Show abstract

The ability to selectively attend to one talker in the presence of competing talkers is crucial to communication. Here we investigate whether cognitive deficits in the absences of hearing loss can impair speech perception. We tested typical hearing, neurodivergent adolescents/adults with autism spectrum disorder, fetal alcohol spectrum disorder, and an age- and sex-matched neurotypical group. We found a strong correlation between IQ and speech perception, with individuals with lower IQ scores having worse speech thresholds. These results demonstrate that deficits in cognitive ability, despite intact peripheral encoding, can impair listening under complex conditions. These findings have important implications for conceptual models of speech perception and for audiological services to improve communication in real-world environments for neurodivergent individuals.

13
Cross-Linguistic Analysis of Speech Markers: Insights from English, Chinese, and Italian Speakers

Santi, G. C.; Catricala, E.; Kwan, S.; Wong, A.; Ezzes, Z.; Wauters, L.; Esposito, V.; Conca, F.; Gibbons, D.; Fernandez, E.; Santos-Santos, M. A.; Chen, T.-F.; Kwan-Chen, L. L.-Y.; Lo, R. R.; Tsoh, J.; Lung-Tat Chen, A.; Garcia, A. M.; de Leon, J.; Miller, Z.; Vonk, J. M. J.; Bruffaerts, R.; Grasso, S. M.; Allen, I. E.; Cappa, S. F.; Gorno-Tempini, M.-L.; Tee, B. L.

2024-10-16 neurology 10.1101/2024.10.15.24314191 medRxiv
Top 0.1%
12.0%
Show abstract

Cross-linguistic studies with healthy individuals are vital, as they can reveal typologically common and different patterns while providing tailored benchmarks for patient studies. Nevertheless, cross-linguistic differences in narrative speech production, particularly among speakers of languages belonging to distinct language families, have been inadequately investigated. Using a picture description task, we analyze cross-linguistic variations in connected speech production across three linguistically diverse groups of cognitively normal participants--English, Chinese (Mandarin and Cantonese), and Italian speakers. We extracted 28 linguistic features, encompassing phonological, lexico-semantic, morpho-syntactic, and discourse/pragmatic domains. We utilized a semi-automated approach with Computerized Language ANalysis (CLAN) to compare the frequency of production of various linguistic features across the three language groups. Our findings revealed distinct proportional differences in linguistic feature usage among English, Chinese, and Italian speakers. Specifically, we found a reduced production of prepositions, conjunctions, and pronouns, and increased adverb use in the Chinese-speakers compared to the other two languages. Furthermore, English participants produced a higher proportion of prepositions, while Italian speakers produced significantly more conjunctions and empty pauses than the other groups. These findings demonstrate that the frequency of specific linguistic phenomena varies across languages, even when using the same harmonized task. This underscores the critical need to develop linguistically tailored language assessment tools and to identify speech markers that are appropriate for aphasia patients across different languages.

14
Semantic and phonetic markers in schizophrenia-spectrum disorders; a combinatory machine learning approach

Voppel, A.; de Boer, J.; Brederoo, S.; Schnack, H.; Sommer, I. e. c.

2022-07-15 psychiatry and clinical psychology 10.1101/2022.07.13.22277577 medRxiv
Top 0.1%
11.9%
Show abstract

IntroductionSpeech is a promising marker for schizophrenia-spectrum disorder diagnosis, as it closely reflects symptoms. Previous approaches have made use of different feature domains of speech in classification, including semantic and phonetic features. However, an examination of the relative contribution and accuracy per domain remains an area of active investigation. Here, we examine these domains (i.e. phonetic and semantic) separately and in combination. MethodsUsing a semi-structured interview with neutral topics, speech of 94 schizophrenia-spectrum subjects (SSD) and 73 healthy controls (HC) was recorded. Phonetic features were extracted using a standardized feature set, and transcribed interviews were used to assess word connectedness using a word2vec model. Separate cross-validated random forest classifiers were trained on each feature domain. A third, combinatory classifier was used to combine features from both domains. ResultsThe phonetic domain random forest achieved 81% accuracy in classifying SSD from HC. For the semantic domain, the classifier reached an accuracy of 80% with a sparse set of features with 10-fold cross-validation. Joining features from the domains, the combined classifier reached 85% accuracy, significantly improving on models trained on separate domains. Top features were fragmented speech for phonetic and variance of connectedness for semantic, with both being the top features for the combined classifier. DiscussionBoth semantic and phonetic domains achieved similar results compared with previous research. Combining these features shows the relative value of each domain, as well as the increased classification performance from implementing features from multiple domains. Explainability of models and their feature importance is a requirement for future clinical applications.

15
Automated Phonological Error Scoring for Children with Language and Hearing Impairment

Sundstrom, S.; Themistocleous, C.

2024-09-04 neurology 10.1101/2024.09.04.24313011 medRxiv
Top 0.1%
11.9%
Show abstract

PurposePhonological production impairments are prevalent in children with developmental language disorder (DLD) and hearing impairment (HI). This study aims to quantify and compare phonological errors in Swedish-speaking children using a novel automated assessment tool and provide an automatic machine learning classification algorithm of children with DLD and HI to age-matched controls based on phonological errors. Methods72 Swedish-speaking children (29 with DLD, 14 with HI, and 29 typically developing) participated. Phonological production was elicited using a 72-item confrontation naming task. A novel tool was developed to calculate a composite phonological error score and specific scores for different phonological errors (deletions, insertions, substitutions, and transpositions) from written speech productions. This tool leverages the International Phonetic Alphabet (IPA) and a form of the Normalized Damerau-Levenshtein Distance for accurate error analysis. ResultsThe composite score successfully differentiated between children with DLD and typically developing children, highlighting its sensitivity in detecting phonological impairment. Machine learning models can accurately differentiate between children with and without language disorders. However, children with DLD and HI differed in the phonemic deletion errors, which suggests that their phonemic production is relatively similar. ConclusionsChildren with DLD and HI exhibit significantly higher phonological error rates compared to typically developing peers. Children with HI and DLC are comparably impaired in phonology (as manifested by the compositive phonological score). These findings highlight the potential of machine learning for early identification and targeted intervention in language disorders, improving outcomes for affected children and demonstrated the potential of a multilingual tool for scoring phonological errors.

16
Towards an extended classification of noise-distortion preferences by modeling longitudinal dynamics of listening choices

Angonese, G.; Buhl, M.; Goesswein, J. A.; Kollmeier, B.; Hildebrandt, A.

2024-10-27 psychiatry and clinical psychology 10.1101/2024.10.25.24316092 medRxiv
Top 0.1%
10.9%
Show abstract

Individuals have different preferences for setting hearing aid (HA) algorithms that reduce ambient noise but introduce signal distortions. "Noise haters" prefer greater noise reduction, even at the expense of signal quality. "Distortion haters" accept higher noise levels to avoid signal distortion. These preferences were assumed to be stable over time, and individuals were classified solely on the basis of these stable, trait scores. However, the question remains as to how stable individual listening preferences are and whether day-to-day state-related variability needs to be considered as a further criterion for classification. We designed a mobile task to measure noise-distortion preferences over two weeks in an ecological momentary assessment study with N = 185 (106 f, Mage = 63.1, SDage = 6.5) unaided individuals with subjective hearing difficulties. Latent State-Trait Autoregressive (LST-AR) modeling was used to evaluate stability and dynamics of individual listening preferences. The analysis revealed a significant amount of state-related variance. The model has been extended to a mixture LST-AR model for data-driven classification, taking into account trait and state components of listening preferences. In addition to successful identification of noise haters, distortion haters and a third intermediate class based on longitudinal, outside of the lab data, we further differentiated individuals with different degrees of variability in listening preferences. It follows that individualisation of HA fitting could be improved by assessing individual preferences along the noise-distortion trade-off, and the day-to-day variability of these preferences needs to be taken into account for some individuals more than others.

17
Automated Macrolinguistic Discourse Analysis for Transdiagnostic Detection of Language Impairments

Lee, S. H.; Wang, S.; Varkanitsa, M.; Kiran, S.

2026-05-21 neurology 10.64898/2026.05.19.26353614 medRxiv
Top 0.1%
10.7%
Show abstract

Macrolinguistic discourse analysis offers valuable insight into how patients with neurogenic communication disorders organize and produce informative speech, yet it remains a largely manual and labor-intensive process. We report an automated pipeline for macrolinguistic discourse analysis for individuals with aphasia and dementia that integrates automatic speech recognition (ASR), utterance segmentation, sentence-level embeddings, centroid-based main-concept matching, and rule-based coherence error classification. These algorithms were applied to Cinderella story retellings from 309 participants (113 controls, 102 post-stroke aphasia (PWA), and 94 dementia). The algorithm reliably identified main concepts (83% accuracy against human labels) and derived interpretable features such as semantic distance to a main concept centroid, main concept coverage, and coherence error rates. Crucially, diagnostic classification results showed that logistic-regression classifiers trained on 10 macrolinguistic features distinguished aphasia from controls with high accuracy (AUC {approx} 0.94) but showed weaker separation for dementia (controls vs dementia AUC {approx} 0.66; aphasia vs dementia AUC {approx} 0.58). Semantic distance to the centroid emerged as a robust, informative predictor for diagnostic classification, demonstrating that the ability to produce narrative-aligned speech is clinically important. The automated pipeline enables scalable macrolinguistic discourse analysis that could support screening and longitudinal monitoring of discourse impairments across neurogenic populations.

18
Automated transcription in primary progressive aphasia: Accuracy and effects on classification

Clarke, N.; Morin, B.; Bedetti, C.; Bogley, R.; Pellerin, S.; Houze, B.; Ramkrishnan, S.; Ezzes, Z.; Miller, Z.; Gorno Tempini, M. L.; Vonk, J. M. J.; Brambati, S. M.

2026-02-26 neurology 10.64898/2026.02.24.26346981 medRxiv
Top 0.1%
10.7%
Show abstract

INTRODUCTIONConnected speech analyses can help characterize linguistic impairments in primary progressive aphasia (PPA) and classify variants, however, manual transcription of speech samples is time-consuming and expensive. Automated speech recognition (ASR) may be efficacious for transcribing PPA speech. METHODSTranscripts of picture descriptions (109 PPA, 32 healthy controls (HC)) were generated using a manual, automated (Whisper) or semi-automated approach including a quality control (QC) step. We evaluated transcript accuracy, the reliability of ASR-derived linguistic features, and classification performance. RESULTSWhisper demonstrated lowest error rates for HC, followed by semantic, logopenic and non-fluent PPA variants. Errors correlated with overall disease severity for semantic and logopenic variants. QC of Whisper outputs reduced errors and improved the reliability of linguistic features. Overall, ASR-derived features achieved better classification performance than manual transcription features. DISCUSSIONResults support the use of off-the-shelf ASR for scalable, cost-efficient transcription of PPA speech and classification.

19
Impaired temporal prediction mechanisms in dyslexia

Bonnet, P. A.; Tillmann, B.; Chettih, E.; Bedoin, N.; Kosem, A.

2026-01-17 neuroscience 10.64898/2026.01.16.699956 medRxiv
Top 0.1%
10.3%
Show abstract

Effective speech analysis involves deconstructing the acoustic signal into identifiable linguistic units, which depends on the ability to recognize and anticipate temporal patterns within the speech stream. However, these processes may be less efficient in individuals with dyslexia. This study investigated the effects of temporal context and related temporal predictions in dyslexic adult participants and matched control participants, using an auditory oddball task with non-verbal stimuli. Pure tones were presented in sequences, and participants were requested to discriminate the pitch of target stimuli. The temporal intervals between the sounds varied in regularity across the sequences, thereby creating contexts with different levels of temporal predictability. At the end of each sequence, participants were prompted to evaluate the perceived rhythmicity of the sequence and to assess their own performance in the auditory discrimination task. Dyslexic participants demonstrated overall lower accuracy in discriminating target sounds than controls. They also showed reduced influence of the temporal context of the sequences on response times, while controls responded faster in sequences that were temporally more regular and predictable. Additionally, individuals with dyslexia perceived the rhythmicity of sound sequences less accurately, overestimating the temporal regularity in irregular sequences and underestimating it in regular sequences. They also reported lower overall confidence in their ability to perform the task compared to control participants. Altogether, these findings provide converging evidence for altered temporal prediction abilities in dyslexia, which may impact auditory perception and then impair language processing.

20
Iceberg or cut off - how adults who stutter articulate fluent-sounding utterances

Leha, A.; Dickhut, S.; Primassin, A.; Korzeczek, A.; Joseph, A. A.; Paulus, W.; Frahm, J.; Sommer, M.

2020-04-17 neuroscience 10.1101/2020.04.15.042432 medRxiv
Top 0.1%
9.7%
Show abstract

Whether fluent-sounding utterances of adults who stutter (AWS) are normally articulated is unclear. We asked 15 AWS and 17 matched adults who do not stutter (ANS) to utter the pseudoword "natscheitideut" 15 times in a 3 T MRI scanner while recording real-time MRI videos at 55 frames per per second in a mid-sagittal plane. All stuttered or otherwise dysfluent runs were discarded. We used sophisticated analyses to model the movement of the tip of the tongue, lips and velum. We observed reproducible movement patterns of the inner and outer articulators which were similar in both groups. Speech duration was similar in both groups and decreased over repetitions, more so in ANS than in AWS. The variability of the movement patterns of tongue, lips and velum decreased over repetitions. The extent of variability decrease was similar in both groups. Across all participants, this repetition effect on movement variability for the lips and the tip of the tongue was less pronounced in severely as compared to mildly stuttering individuals. We conclude that there is no major difference in the movement patterns of a fluent-sounding utterance in both groups. This encourages studies looking at state rather than trait markers of speech dysfluency.